Microsoft2026-10-03 02:32:34Microsoft rolls out three audio models for voice agents, with speech generation latency as low as 45msMicrosoft has introduced three audio models in one move: MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Together, they cover the full voice-agent pipeline, from real-time speech-to-text to text-to-speech generation, with a focus on low-latency use cases. MAI-Transcribe-2-Streaming is built for live transcription and can keep producing text before a speaker finishes a sentence. It supports 60 languages and automatic language detection. According to Artificial Analysis’ streaming speech transcription ranking, the model placed first in both final transcription accuracy and first partial transcription accuracy, posting a 2.5% word error rate. Its final text is returned about 0.13 seconds after speech end is detected, compared with 2.7% and 0.49 seconds for Grok Voice Transcribe 2.0. On the speech generation side, MAI-Voice-2.1 supports 23 languages and allows a single voice to switch across languages. The Flash version is tuned for speed, with inference at about 45ms and pricing of $15 per million characters, while the standard version runs at about 550ms and costs $22 per million characters. Both text-to-speech models can match a voice using a short reference audio sample.20
Microsoft AI2026-10-02 09:31:52Microsoft AI rolls out two speech models for live transcription and voice agentsMicrosoft AI has introduced two new models, MAI-Transcribe-2-Streaming and MAI-TTS-2, aimed at improving capabilities for voice agents. According to a Techub News item citing The Decoder, the first model focuses on real-time speech-to-text transcription, while the second converts text into natural-sounding speech. Both were tuned for voice-agent use cases rather than presented as general-purpose releases. The announcement centers on two core functions that sit at the heart of spoken AI systems: converting live audio into text with low latency, and generating spoken responses from written input. In the brief release, Microsoft AI framed the pair as tools designed to strengthen speech-based agent experiences. No additional technical specifications, launch regions, pricing details, or deployment timelines were disclosed in the source item.50
ElevenLabs2026-09-29 10:37:02ElevenLabs launches v4 and Turbo, topping TTS benchmark ahead of Gemini and QwenElevenLabs has released its new voice models, Eleven v4 and the real-time focused Eleven v4 Turbo, with the flagship model moving to the top of the latest Artificial Analysis text-to-speech blind test ranking. Eleven v4 is listed at 1315 Elo, ahead of Cartesia Sonic 3.6 at 1275, Gemini 3.8 Flash TTS at 1267, and Qwen-Audio-3.0-TTS-Plus at 1258. The company’s previous-generation Eleven v3 currently stands at 1169 Elo, placing 17th in the same ranking. The new release centers on tighter control over tone, pacing, emotion, and sound effects. According to the source material, users can either specify those parameters directly or describe in natural language how a sentence should be spoken. ElevenLabs also expanded language coverage from more than 70 languages to 90-plus languages, while Instant Voice Clone requires about 10 seconds of audio. The company said the update also improves voice consistency in long-form generation and multi-speaker dialogue. For latency, official documentation shows median inference time for v3 Conversational at about 280ms, while v4 Turbo cuts that to roughly 100ms for real-time voice agent use cases.160
OpenAI2026-09-15 10:52:26GPT-Live-1 tops speech agent ranking after adding Astra, edging past GrokArtificial Analysis’ latest benchmark shows that OpenAI’s GPT-Live-1 moved to the top of the Speech to Speech Index after connecting GPT-6 Astra as its backend model. The system scored 81.5, slightly above Grok Voice Think Fast 2.0 High at 81.3. According to the benchmark description, GPT-Live-1 handles real-time listening and speaking, while more complex reasoning and tool-use tasks are handed off to Astra. The combined setup did not take first place in speech reasoning alone. Still, it ranked No. 1 in an agent evaluation focused more on practical task execution, posting 67.9%. That result lifted its overall score to the top of the leaderboard. The ranking highlights how model orchestration, rather than a single-model design, can improve end-to-end performance in voice agent testing.610